Papers with language modelling

35 papers
Keep Learning: Self-supervised Meta-learning for Learning from Inference (2021.eacl-main)

Copied to clipboard

Challenge: A common approach to improve performance of machine learning algorithms involves self-supervised learning on large unlabeled data before fine-tuning on downstream tasks.
Approach: They propose to use model's own class-balanced predictions to back-propagate the loss from the model''s class-balancing predictions (pseudo-labels) this method improves performance of standard backbones such as BERT, Electra, and ResNet-50 on a wide variety of tasks, including question answering on SQuAD and NewsQA .
Outcome: The proposed method outperforms previous approaches on a wide variety of tasks including question answering on SQuAD and NewsQA, benchmark task SuperGLUE, conversation response selection on Ubuntu Dialog corpus v2.0, and image classification on MNIST and ImageNet.
Generating Text through Adversarial Training Using Skip-Thought Vectors (N19-3)

Copied to clipboard

Challenge: Existing approaches to use word embeddings for text generation have been limited.
Approach: They propose to use GANs with word embeddings to reproduce writing style in text . they use a sentence embeddable vector to model people's way of expression .
Outcome: The proposed model outperforms baseline text generation networks across several metrics including BLEU-n, METEOR and ROUGE.
Enhancing Transformers with Gradient Boosted Decision Trees for NLI Fine-Tuning (2021.findings-acl)

Copied to clipboard

Challenge: Recent advances in transfer learning have brought significant improvements to many natural language processing tasks.
Approach: They propose a method of fitting a GBDT head on the features computed during finetuning to increase performance without additional computation by the neural network.
Outcome: The proposed method improves on several NLI datasets using a strong baseline model (RoBERTa-large) with MNLI pretraining.
Graph-Induced Syntactic-Semantic Spaces in Transformer-Based Variational AutoEncoders (2024.findings-naacl)

Copied to clipboard

Challenge: Existing studies on syntactic injection in Variational AutoEncoders (VAEs) are limited to LSTM-based VAEs.
Approach: They propose to use latent space separation techniques to inject syntactic information into Variational AutoEncoders (VAEs) using graph-based models.
Outcome: The proposed end-to-end VAE architecture can improve the organisation of the latent space, alleviating the information loss occurring in standard VAE setups, and resulting in enhanced performances on language modelling and downstream generation tasks.
Re-framing Incremental Deep Language Models for Dialogue Processing with Multi-task Learning (2020.coling-main)

Copied to clipboard

Challenge: Using a multi-task learning framework, we train a universal incremental dialogue processing model with four tasks of disfluency detection, language modelling, part-of-speech tagging and utterance segmentation in a simple deep recurrent setting.
Approach: They propose a multi-task learning framework to train a universal incremental dialogue processing model with four tasks of disfluency detection, language modelling, part-of-speech tagging and utterance segmentation in a simple deep recurrent setting.
Outcome: The proposed model outperforms individual tasks and delivers competitive performance.
Cascaded Semantic and Positional Self-Attention Network for Document Classification (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to document classification combine semantic information with positional information (word orders) . document classification is one of the fundamental problems in natural language processing .
Approach: They propose a new architecture to combine semantic and positional information using a semantic self-attention layer cascaded with Bi-LSTM.
Outcome: The proposed model can exploit the interaction between semantics and word positions in a more interpretable and adaptive manner while preserving a compact model size and high convergence rate.
T3L: Translate-and-Test Transfer Learning for Cross-Lingual Text Classification (2023.tacl-1)

Copied to clipboard

Challenge: Existing approaches to cross-lingual text classification leverage text classifiers trained in a high-resource language to perform text classification in other languages with no or minimal fine-tuning.
Approach: They propose to combine a neural machine translator and a text classifier trained in a high-resource language to perform text classification in other languages with no or minimal fine-tuning.
Outcome: The proposed approach significantly improves over a baseline approach.
Efficiently and Thoroughly Anonymizing a Transformer Language Model for Dutch Electronic Health Records: a Two-Step Method (2022.lrec-1)

Copied to clipboard

Challenge: Neural Networks (NNs) are used to model large amounts of data, such as text data, and have shown to be very useful for language modelling.
Approach: They propose to use a Dutch language model for hospital notes to anonymize a model trained on large amounts of data and publish it online.
Outcome: The proposed method predicts a name-like token 0.2% of the time, compared to the original training data.
Can Activation Steering Generalize Across Languages? A Study on Syllogistic Reasoning in Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Prior work has focused on activation steering for Large Language Models (LLMs) this technique can be used to improve reasoning accuracy and transferability across languages.
Approach: They propose to use activation steering to steer models towards a cross-lingual reasoning space.
Outcome: The proposed techniques generalise well to multilingual datasets while minimizing language modelling performance.
Surgical Feature-Space Decomposition of LLMs: Why, When and How? (2024.acl-long)

Copied to clipboard

Challenge: Low-rank approximations of the weight and feature space can enhance the performance of large language models.
Approach: They propose to use weight and feature space decomposition to improve LLM performance . they also extend their investigation to the implications of low-rank approximations on model bias .
Outcome: The proposed low-rank approximations can improve performance of large language models . the authors show that the approximate can improve generalization and inference performance .
Language Modelling as a Multi-Task Problem (2021.eacl-main)

Copied to clipboard

Challenge: Using multitask learning, humans are optimising their behaviour towards a multitude of objectives to reach their goals in dayto-day life.
Approach: They propose to study language modelling as a multi-task problem by examining the generalisation behaviour of language models as they learn the linguistic concept of Negative Polarity Items.
Outcome: The proposed model is able to learn the linguistic concept of Negative Polarity Items (NPIs) and is a multi-task learning model.
Improving Variational Autoencoder for Text Modelling with Timestep-Wise Regularisation (2020.coling-main)

Copied to clipboard

Challenge: Variational Autoencoders (VAEs) have been widely used in text modelling but posterior collapse is a problem when RNN-based models are employed.
Approach: They propose a timestep-wise regularisation VAE architecture which can effectively avoid posterior collapse when used in text modelling.
Outcome: The proposed model avoids posterior collapse and can be applied to any RNN-based VAE model.
Multi-task Learning of Negation and Speculation for Targeted Sentiment Classification (2021.naacl-main)

Copied to clipboard

Challenge: Currently, most work on targeted sentiment analysis is focused on improving the overall results.
Approach: They propose a multi-task learning method to incorporate information from syntactic and semantic auxiliary tasks to create English-language models that are more robust to linguistic phenomena.
Outcome: The proposed method improves on negation and speculation datasets but there is room for improvement.
Towards Zero-shot Language Modeling (D19-1)

Copied to clipboard

Challenge: a number of natural questions have been asked about the inductive biases of neural networks on core NLP tasks.
Approach: They construct an informative prior for held-out languages on a task of character-level, open-vocabulary language modelling.
Outcome: The proposed model outperforms baseline models with an uninformative prior in both zero-shot and few-shot settings, showing that it is imbued with universal linguistic knowledge.
A Second Wave of UD Hebrew Treebanking and Cross-Domain Parsing (2022.emnlp-main)

Copied to clipboard

Challenge: Foundational Hebrew NLP tasks have relied on various versions of the Hebrew Treebank . however, the data in the HTB is now over 30 years old and does not cover many aspects of contemporary Hebrew on the web.
Approach: They propose to use Hebrew Wikipedia to stratify the text from a UD treebank.
Outcome: The proposed treebank is based on a single-source newswire corpus selected from Hebrew Wikipedia.
Evaluating the Impact of Sub-word Information and Cross-lingual Word Embeddings on Mi’kmaq Language Modelling (2020.lrec-1)

Copied to clipboard

Challenge: Mi'kmaq is an Indigenous language spoken primarily in Eastern Canada.
Approach: They consider n-gram and RNN language models for Mi'kmaq and use them to investigate their performance.
Outcome: The proposed model performs better than word-level models, but does not improve over word-based models.
No Data to Crawl? Monolingual Corpus Creation from PDF Files of Truly low-Resource Languages in Peru (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods for extracting text from PDF files are expensive and limited by the absence of web content of endangered languages.
Approach: They propose a method for creating monolingual corpora for four endangered languages . they use a PDF file format with multilingual sentences and noisy pages .
Outcome: The proposed method allows the creation of clean corpora for the four languages, a key resource for natural language processing tasks nowadays.
Max-Margin Incremental CCG Parsing (2020.acl-main)

Copied to clipboard

Challenge: a new incremental parser reduces the number of beam search violations and minimises the biggest violation.
Approach: They propose to use beam search optimisation to minimise all beam search violations instead of minimising only the biggest violation.
Outcome: The proposed parser outperforms existing non-incremental parsers and minimises all beam search violations instead of minimising the biggest violation.
Can Large Language Models Learn Independent Causal Mechanisms? (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) perform poorly on complex reasoning tasks, such as abstract, causal, or logical reasoning.
Approach: They propose to use two concepts from causality to learn ICMs within LLMs to improve out-of-distribution performance on abstract and causal reasoning tasks.
Outcome: The proposed model outperforms existing models on abstract and causal reasoning tasks and is more robust to fine-tuning.
Analysing The Impact of Sequence Composition on Language Model Pre-Training (2024.acl-long)

Copied to clipboard

Challenge: Existing studies show that pretraining sequence composition strategy can lead to distracting information from previous documents.
Approach: They propose to use a sequence construction method to concatenate documents into fixed-length sequences to compute the likelihood of each token given its context.
Outcome: The proposed method can improve in-context learning, knowledge memorisation and context utilisation without sacrificing efficiency.
Rollenwechsel-English: a large-scale semantic role corpus (L18-1)

Copied to clipboard

Challenge: The Rollenwechsel-English corpus is a large corpus of automatically-labelled semantic frames extracted from the ukWaC corpus and BNC using Propbank roles.
Approach: They present a large corpus of automatically-labelled semantic frames extracted from ukWaC and BNC using Propbank roles.
Outcome: The rollenwechsel-English corpus is a large corpus of automatically-labelled semantic frames extracted from the ukWaC corpus and BNC using Propbank roles.
Subword Segmental Language Modelling for Nguni Languages (2022.findings-emnlp)

Copied to clipboard

Challenge: Subword segmentation is a standard practice in NLP, but is viewed as a preprocessing step for low-resource languages with complex morphologies.
Approach: They propose a subword segmental language model that learns how to segment words while being trained for autoregressive language modelling.
Outcome: The proposed model outperforms existing models on unsupervised morphological segmentation and outperfies standard subword segmenters on all 4 languages.
Training Neural Response Selection for Task-Oriented Dialogue Systems (P19-1)

Copied to clipboard

Challenge: Despite their popularity, retrieval-based models have had modest impact on task-oriented dialogue systems . main obstacle to their application is the low-data regime of most task-orientated dialogue tasks . e-commerce, banking, and other domains are applications of retrieval models .
Approach: They propose a method which pretrains a retrieval-based model on large general-domain conversational corpora and fine-tunes it for the target dialogue domain.
Outcome: The proposed method is evaluated on five diverse domains, ranging from e-commerce to banking.
Latvian National Corpora Collection – Korpuss.lv (2022.lrec-1)

Copied to clipboard

Challenge: Latvian National Corpora Collection (LNCC) is a multi-institutional and multi-project effort supporting the Latvian language research and language modelling.
Approach: They propose to use Latvian corpora for linguistic research and language modelling.
Outcome: LNCC is a multi-institutional and multi-project effort supported by the Digital Humanities and Language Technology communities in Latvia.
Multilingual Multi-Figurative Language Detection (2023.findings-acl)

Copied to clipboard

Challenge: Figures of speech help people express abstract concepts and emotions, but it's understudied in a multilingual setting and when considering more than one figure of speech at the same time.
Approach: They propose a framework for sentence-level figurative language detection based on template-based prompt learning and use it to unify multiple detection tasks that are interrelated across multiple figures of speech and languages.
Outcome: The proposed framework outperforms baselines and may serve as blueprint for the joint modelling of other interrelated tasks.
EmbedTextNet: Dimension Reduction with Weighted Reconstruction and Correlation Losses for Efficient Text Embedding (2023.findings-acl)

Copied to clipboard

Challenge: EmbedTextNet is a light add-on network that can be appended to an arbitrary language model to generate a compact embedding without requiring any changes in its architecture or training procedure.
Approach: They propose an add-on network that can be appended to an arbitrary language model to generate a compact embedding without requiring any changes in its architecture or training procedure.
Outcome: The proposed network can be appended to an arbitrary language model to generate a compact embedding without any changes in its architecture or training procedure.
Improving Code-switched ASR with Linguistic Information (2022.coling-1)

Copied to clipboard

Challenge: Existing studies on code-switching have been limited to the individual languages, but the results are promising.
Approach: They propose to apply linguistic theories to generate more realistic code-switching text, which is needed for language modelling in ASR.
Outcome: The proposed system improves 2% on English-Spanish code-switching . Equivalence Constraint theory and part-of-speech labelling are particularly helpful for text generation and bring 2% improvement to ASR performance.
Effective Estimation of Deep Generative Language Models (2020.acl-main)

Copied to clipboard

Challenge: Existing techniques for parameterisation of probabilistic models by deep neural networks are difficult to use in language modelling due to posterior collapse.
Approach: They propose to use variational auto-encoder to estimate probabilistic models of language by deep neural networks.
Outcome: The proposed model performs reasonably well given enough resources, but a favourite can be named based on convenience.
Towards Language Technology for Mi’kmaq (L18-1)

Copied to clipboard

Challenge: Mi'kmaq is a polysynthetic Indigenous language spoken primarily in Eastern Canada .
Approach: They construct and analyze a web corpus of Mi'kmaq and evaluate several approaches to language modelling . they argue that natural language processing could aid efforts to preserve Indigenous languages .
Outcome: The proposed language model is based on a web corpus of Mi'kmaq . the model is well-suited to morphologically-rich languages, the authors argue .
Do Transformers Need Deep Long-Range Memory? (2020.acl-main)

Copied to clipboard

Challenge: Deep attention models have advanced the modelling of sequential data across many domains.
Approach: They propose to use a Transformer augmented with a long-range memory to model sequential data across many domains.
Outcome: The Transformer-XL has a long-range memory at every layer of the network, rendering its state thousands of times larger than RNN predecessors.
Classist Tools: Social Class Correlates with Performance in NLP (2024.acl-long)

Copied to clipboard

Challenge: despite growing concerns surrounding fairness and bias in NLP, there is a dearth of studies delving into the effects it may have on NLP systems.
Approach: They argue that NLP systems’ performance is affected by speakers’ SES, potentially disadvantaging less-privileged socioeconomic groups.
Outcome: The proposed model shows that NLP systems perform better on tasks with social class, ethnicity and geographical variation than those without social class.
SliceMoE: Routing Embedding Slices Instead of Tokens for Fine-Grained and Balanced Transformer Scaling (2025.emnlp-main)

Copied to clipboard

Challenge: Token-level routing assigns an entire semantic spectrum to each expert, creating capacity bottlenecks, load-balancing pathologies, and limited specialisation.
Approach: They propose an architecture that routes contiguous slices of a token’s hidden vector and a lightweight shared router predicts the top-k experts.
Outcome: The proposed architecture achieves 1.7x faster inference than dense baselines, 12–18% lower perplexity than parameter-matched token-MoE, and improved expert balance.
SuperTweetEval: A Challenging, Unified and Heterogeneous Benchmark for Social Media NLP Research (2023.findings-emnlp)

Copied to clipboard

Challenge: specialised language models (LMs) have shown to exhibit lower perplexity and higher downstream performance across the board.
Approach: They propose a benchmark for NLP evaluation in social media, SuperTweetEval.
Outcome: The proposed benchmark shows that social media models perform better when compared to general-purpose models, metrics and benchmarks.
A Simple and Effective L_2 Norm-Based Strategy for KV Cache Compression (2024.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to reduce the KV cache size involve fine-tuning the model to learn a compression strategy or leveraging attention scores to reduce sequence length.
Approach: They find a correlation between the L2 norm and attention scores over cached KV pairs . they compress the KV cache based on the L1 norm of key embeddings .
Outcome: The proposed approach reduces the KV cache size by 50% on language modelling and needle-in-a-haystack tasks and 90% on passkey retrieval tasks without losing accuracy.
Causal Estimation of Tokenisation Bias (2025.acl-long)

Copied to clipboard

Challenge: Modern language models define probabilities over character-strings, but in practice, it does . Ideally, the choice of the tokeniser should not affect the probability assigned to the underlying character- string.
Approach: They quantify a type of tokenisation bias by framing it as a causal effect and estimating it using the regression discontinuity design.
Outcome: The proposed model can estimate tokenisation bias by comparing subwords around arbitrary cutoff points.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations